Papers with dataset curation

7 papers
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing evaluation methods do not explicitly measure this distinction, hindering effective dataset curation and real-world focused model development.
Approach: They introduce a region-based score to quantify a dataset's reliance on global versus local visual information.
Outcome: The proposed model-based score systematically compares model performance on image patches versus full images to determine if tasks require holistic image understanding or can be solved with partial or localized visual cues.
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for video temporal grounding suffer from limited temporal awareness and poor generalization.
Approach: They propose a two-stage training framework that integrates supervised fine-tuning with reinforcement learning to improve both the accuracy and robustness of VTG models.
Outcome: The proposed training framework outperforms existing models on multiple benchmarks on open-domain and challenging scenarios.
On the Feasibility of In-Context Probing for Data Attribution (2025.findings-naacl)

Copied to clipboard

Challenge: In-context probing (ICP) can be used to identify training data that contributes to model outputs, but many data attribution methods, such as influence functions, use model gradients and are computationally expensive.
Approach: They propose to use in-context probing (ICP) to proxy for gradient-based data attribution for data selection under conditions contingent on data similarity.
Outcome: The proposed method can be used to identify training data that contribute to model outputs and fine tune models on training data.
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters (2026.acl-long)

Copied to clipboard

Challenge: Existing research lacks systematic analysis of the applicability and methodology of cross-modal skill injection.
Approach: They investigate the applicability and methodology of cross-modal skill injection by integrating a domain-expert LLM into a VLM.
Outcome: The proposed method enables transfer of domain-specific expertise from Large Language Models (LLMs) to VLMs without incurring additional training data requirements or significant computational overhead.
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models are fine-tuned on diverse datasets to develop a range of skills . each skill has unique characteristics, and datasets are heterogeneous and imbalanced . a general, model-agnostic, reinforcement learning framework is proposed to optimize data usage .
Approach: They propose a general, model-agnostic, reinforcement learning framework that optimizes data usage automatically during the fine-tuning process.
Outcome: The proposed framework optimizes data usage automatically during the fine-tuning process.
Towards a new research agenda for multimodal enterprise document understanding: What are we missing? (2024.findings-acl)

Copied to clipboard

Challenge: In this paper, we discuss the limitations of multimodal document understanding models in enterprise settings.
Approach: They propose a research agenda that is aimed at driving the field towards higher impact in enterprise applications.
Outcome: The proposed research agenda is aimed at driving the field towards higher impact in enterprise applications.
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency (2025.emnlp-main)

Copied to clipboard

Challenge: Current dataset curation and bias assessment practices lack transparency . current approaches lack a thorough understanding of how data characteristics influence model behavior .
Approach: They propose a comprehensive bias evaluation framework that integrates general benchmarks with a healthcare-specific methodology to probe for biases in a sensitive healthcare context.
Outcome: The proposed approach to bias evaluation leverages established benchmarks and a healthcare-specific methodology.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations